Back

Epigenetics & Chromatin

Springer Science and Business Media LLC

Preprints posted in the last 90 days, ranked by how well they match Epigenetics & Chromatin's content profile, based on 42 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
Histone modifications analysis reveals enhancers reprogramming during maternal-to-zygotic transition

Hu, K.; Wang, C.; Fang, D.; Lu, J.; Meng, X.; Chen, L.; Yao, Y.; Guo, J.; Khan, S.; Li, W.; Wang, Y.; li, Y.; Chen, H.; Xu, J.

2026-05-09 developmental biology 10.64898/2026.05.06.723106 medRxiv
Top 0.1%
5.1%
Show abstract

Enhancers are key epigenetic regulatory elements that orchestrate spatiotemporal gene expression and are critical in mammalian development, gene regulation, and disease. Histone modifications such as H3K4me1 (a canonical enhancer mark) and H3K27ac (which distinguishes active enhancers) remain poorly characterized during early mammalian embryogenesis. Using low-input CUT&RUN (Cleavage Under Targets and Release Using Nuclease) with input as low as 50 cells, this study profiles genome-wide H3K4me1 and H3K27ac patterns in mouse oocytes and pre-implantation embryos. Both marks are enriched in distal regions and exhibit distinct sequence preferences and reprogramming dynamics in pre-implantation embryos. H3K27ac is reprogrammed at the 2-cell stage and marks active enhancers, while H3K4me1 is remodeled at the 4-cell stage and co-localizes with H3K27ac, overlapping with accessible chromatin regions. Interestingly, the co-localization of H3K4me1 and H3K27ac is also detected in promoter regions, where they exhibit a mutually exclusive pattern with H3K4me3. Three enhancer types-active (H3K4me1/H3K27ac), primed (H3K4me1), and poised (H3K4me1/H3K27me3)-are dynamically remodeled during maternal-to-zygotic transition (MZT), with active enhancers increasing significantly after zygotic genome activation. Furthermore, genome-wide super-enhancers are identified and mainly enriched in promoters. The differences in gene expression at different stages may be related to the specific motifs enriched by super-enhancers.

2
Modulating Nucleosomal H3 Tail Dynamics with Lysine and Serine Modifications

Adkins, B. J.; Sidlowski, P. F. W.; Jennings, C. E.; Morrison, E. A.

2026-07-03 biophysics 10.64898/2026.06.30.735535 medRxiv
Top 0.1%
4.0%
Show abstract

Nuclear organization is dynamic and originates from the fundamental subunit of chromatin, the nucleosome. Post-translational modification of nucleosomal histones, particularly within intrinsically disordered histone tail regions, provides a dynamic regulatory mechanism of accessibility for chromatin-templated processes. While the epigenomic impacts of lysine acetylation and serine phosphorylation in the histone H3 tail are well-known, how these charge-altering post-translational modifications (PTMs) alter nucleosomal tail conformational dynamics remains incompletely characterized. Given that the functional implications of these PTMs are, at least in part, a consequence of modified nucleosome conformation, systematically cataloging the impact of histone PTMs on nucleosome dynamics provides crucial insight into both baseline cellular activity and epigenetic dysregulation that occurs in disease. Previously, our lab demonstrated that arginine citrullination mimetics lead to regional increases in H3 tail dynamics within nucleosome core particles. Here, we performed nuclear magnetic resonance spin relaxation experiments to investigate the effects of lysine acetylation and serine phosphorylation on H3 tail picosecond-nanosecond (ps-ns) dynamics. Using lysine-to-glutamine and serine-to-glutamate mutations as acetyllysine and phosphoserine mimetics, respectively, we found that these PTMs increase ps-ns conformational dynamics regionally around the PTM site, with a position-dependent effect. Additionally, we show that the type of PTM influences the extent of these increases: in general, the effect of mimetics trends in the order of phosphorylation [&le;] acetylation < citrullination, suggesting a tunable method for altering histone tail dynamics. Taken together, these results illustrate the role of nucleosome conformational dynamics in conveying the effects of epigenomic PTMs, elucidating a mechanism of the histone language.

3
eQTM (expression quantitative trait methylation) Atlas: a comprehensive resource of over 11 million DNA methylation-gene expression associations through across 11 tissues and 4 diseases

Sriram, A.; Kim, S.; Caldino Bohn, R.; Chen, W.; Liu, T.; Yue, M.; Jain, N.; Pierce, B.; Joehanes, R.; Levy, D.; Patin, E.; Quintana-Murci, L.; Park, H. J.; Celedon, J. C.

2026-06-10 genetics 10.64898/2026.06.07.730721 medRxiv
Top 0.1%
3.3%
Show abstract

MotivationEpigenome-wide association studies (EWAS) have identified numerous DNA methylation (DNAm) CpG sites associated with complex traits and diseases, but interpretation of those CpG sites remains challenging because in EWAS, CpGs are mostly linked to nearby genes based only on genomic proximity. Expression quantitative trait methylation (eQTM) analyses connect DNAm CpGs with statistically associated gene expression levels. However, a comprehensive, searchable resource integrating eQTMs across diverse tissues and disease contexts has been lacking. ResultsWe developed the eQTM Atlas, a web-based resource that manually curates more than 11 million DNAm-gene expression associations from eight cohorts, covering 11 tissue types, four broad disease contexts, 173,886 unique CpG probes and 20,231 unique genes. The Atlas supports gene- or CpG-searches by tissue or disease type and finding associated CpG or genes, visualization of cis- and trans-eQTMs through genome browser, heatmap interfaces across various tissues, and cohort-level data downloads. By integrating eQTM results with EWAS resources, the eQTM Atlas enables users to connect disease- or trait-associated CpGs to statistically associated genes rather than relying solely on proximity-based gene annotation, supporting functional interpretation of EWAS findings and generation of disease-specific regulatory hypotheses. Availability and implementationThe eQTM Atlas is freely available at https://shiny.crc.pitt.edu/eqtm_browser/. The web interface is implemented in R Shiny and hosted through the University of Pittsburgh Center for Research Computing (CRC). Source code is available at https://github.com/ads303/eQTM-Atlas.

4
Differential histone tail citrullination by PAD Enzymes observed via NMR spectroscopy

Kowalczyk, A. J.; Morrison, E. A.

2026-05-05 biophysics 10.64898/2026.05.01.722238 medRxiv
Top 0.2%
3.2%
Show abstract

Citrullination is a charge-modifying post-translational modification whereby proteinogenic arginine is converted to the non-coded amino acid citrulline by calcium-activated protein arginine deiminases (PADs; EC 3.5.3.15). The five known PAD enzymes in humans (PADs 1, 2, 3, 4, and 6) are differentially expressed and have distinct targets, including histones. While some PAD histone citrullination sites are known, a comprehensive investigation of all histone tail arginines targeted by catalytically active PADs 1-4 is lacking. Here, we sought to identify PAD citrullination sites in histone tails, both within histone peptides and in reconstituted nucleosomes. Toward this objective, we utilized a real-time 1H-15N NMR spectroscopy-based assay. By monitoring both arginine and citrulline backbone amide peak intensities over time, we identified sites of citrullination in 15N-labeled histone tails within peptides and reconstituted nucleosome core particles. We found that PADs 1, 2, and 4 citrullinate all directly observable histone tail arginines to varying degrees. This is distinct from PAD3, which only moderately citrullinates H2A and H4 arginine residues and does not modify H3 tail arginines. Together, these data suggest a level of histone arginine specificity by each PAD. Furthermore, histone tail citrullination is altered within nucleosomes compared to isolated peptides, which we interpret to reflect changes in conformation and accessibility. We speculate that citrullination increases nucleosomal histone tail dynamics, with implications for crosstalk between sites of histone citrullination and other important sites of regulation by PTMs (including lysines) within and between tails.

5
Smoking drives an epigenetic memory of aberrant hematopoiesis

Breeze, C. E.; Goodney, G.; Wang, H.; Hubbard, A. K.; Lim, J.; Machiela, M. J.; Hoang, T. T.; Richards-Barber, M.; Tran, C.; Tolentino, M.; Hansen, M.; Porecha, R.; Renke, N.; Zhou, W.; Franceschini, N.; Berndt, S. I.; Hofmann, J.; Lee, M.; London, S. J.; Wong, J. Y.

2026-05-21 epidemiology 10.64898/2026.05.14.26353250 medRxiv
Top 0.2%
2.5%
Show abstract

Tobacco smoking induces DNA methylation (DNAm) changes in blood and other tissues, which may influence chronic health outcomes. However, the breadth of smoking-related DNAm changes remains unmapped, offering a space for employing novel technologies. To expand our understanding of smoking impacts on DNAm, we conducted an epigenome-wide association study (EWAS) comparing ever smokers to never smokers, using blood from a multiethnic U.S. study population (n=887). We employed the newly developed Illumina Methylation Screening Array (MSA) covering 269,094 unique sites, including 123,776 CpGs not assayed in previous EWAS. Trans-ethnic meta-analysis identified 152 differentially methylated positions (DMPs) associated with ever-smoking status (n=764); European-specific analysis yielded 129 DMPs (n=674), including 106 overlapping with trans-ethnic analysis. A separate, large-scale replication EWAS (n=2,190) confirmed 91 trans-ethnic and 77 European-specific DMPs. Among our findings, we identified 61 DMPs at CpGs novel to the MSA platform, including near both new and known smoking-associated genes. Most notably, we uncovered a dense cluster of 12 DMPs within a 1117 bp region of ECEL1P1, forming the most long-lasting, persistent smoking-associated DMR ever detected, even among former smokers who quit decades prior. We also detected new signals at AHRR, a well-known locus for smoking-related DNAm changes. eFORGE analysis revealed that detected smoking-associated DNAm changes are predominantly located in hematopoietic stem and progenitor cell (HSPC) DNase I hotspots, aligning with gene set enrichment analyses that highlighted pathways related to hematopoietic stem cell differentiation. Our findings suggest that HSPCs serve as a reservoir for an epigenetic memory of smoking. Additionally, we observed short-term cell-specific smoking-associated DNAm changes in myeloid cells. Our results demonstrate the utility of the MSA in expanding our knowledge of both transient and persistent environmental exposure-associated DNAm changes.

6
Folate-dependent one-carbon metabolism controls meiotic and post-meiotic epigenome remodeling in the male germline

Ikuyo, A.; Fuse, N.; Mori, M.; Hirayama, A.; Yamada, Y.; Nakamura, T.; Sagi, T.; Otsuka, K.; Namekawa, S. H.; Hayashi, Y.; Maezawa, S.

2026-04-28 developmental biology 10.64898/2026.04.23.720283 medRxiv
Top 0.3%
1.9%
Show abstract

Environmental exposures can influence offspring health through epigenetic alterations in the male germline. Folate deficiency, a dietary perturbation that disrupts one-carbon metabolism and S-adenosylmethionine (SAM) production, has been linked to altered histone methylation and developmental abnormalities in offspring. However, when and how folate availability shapes the germline epigenome during spermatogenesis remains unclear. In this study, unbiased metabolomic profiling of spermatogenic cells uncovers stage-specific metabolic remodeling, including downregulation of serine- glycine-one-carbon (SGOC) metabolism in meiotic spermatocytes. Using a post-weaning folate-deficient mouse model, we investigate how folate availability influences germline epigenome establishment during spermatogenesis. Consistent with this metabolic transition, genome-wide chromatin accessibility profiling demonstrates extensive, stage-dependent remodeling under folate-deficient conditions, particularly in meiotic spermatocytes and post-meiotic spermatids. These accessibility changes display cell-type-specific genomic distributions and preferential localization to repressive chromatin compartments in post-meiotic cells. Histone modification analyses further reveal bidirectional redistribution of the active histone mark H3K4me3 in round spermatids. Although genome-wide distribution of the repressive mark H3K27me3 remains largely stable, folate deficiency alters its nuclear organization. Notably, a subset of H3K4me3 alterations established in post-meiotic cells is retained in mature sperm, providing a mechanistic link between paternal metabolic perturbation and the germline epigenome. Together, these findings demonstrate that folate availability shapes germline epigenome establishment through stage-specific metabolic and chromatin remodeling during spermatogenesis, revealing a metabolic basis for paternal environmental effects on the germline epigenome.

7
Average local nucleosome motion remains constant during interphase in living human cells

Nagata, Y.; Iida, S.; Shimazoe, M. A.; Tamura, S.; Nakazato, K.; Shimizu, K.; Hatoyama, Y.; Kanemaki, M.; Maeshima, K.

2026-05-01 cell biology 10.64898/2026.04.29.721002 medRxiv
Top 0.3%
1.8%
Show abstract

BackgroundDynamic chromatin behavior, which is related to chromatin accessibility, plays a critical role in various genome DNA functions such as RNA transcription and DNA replication/repair. Previous studies using highly synchronized cells showed that average local chromatin motion, captured by single-nucleosome imaging and tracking on a second time scale, remained almost constant throughout G1, S, and G2 phases in living human cells, although possible effects of prolonged drug treatments for cell-cycle synchronization could not be excluded. ResultsTo avoid possible effects of prolonged drug treatment, we combined single-nucleosome imaging with Fucci probes to visualize cell-cycle progression through G1, S, and G2. Using HeLa and HCT116 cells expressing H2B-HaloTag and Fucci probes, we found that local nucleosome motion remained similar on average throughout interphase, except for elevated motion in early G1. Transcription inhibition similarly increased nucleosome motion throughout interphase. Local nucleosome motion also increased following replication stress or DNA damage. ConclusionOur findings suggest that near-constant chromatin motion supports housekeeping functions under similar physical conditions during interphase. Our findings also suggest that cells can transiently change chromatin motion to perform ad hoc tasks in response to signals from inside and outside the cell, such as DNA damage.

8
nanoASM: Long-Read Allele-Specific DNA Methylation Profiling Enables Functional Annotation of Regulatory Noncoding Variants in Human Prostate Tissues

Tian, Y.; Wong, J.; McDonnell, S.; Zhong, H.; Wu, L.; Larson, N.; Manley, B. J.; Wang, L.

2026-06-22 genetics 10.64898/2026.06.17.732357 medRxiv
Top 0.3%
1.7%
Show abstract

Long-read nanopore sequencing enables simultaneous detection of germline variation and native DNA base modifications on individual DNA molecules, providing a unique opportunity to investigate allele-specific epigenetic regulation. Here, we performed whole-genome nanopore sequencing on normal and tumor prostate tissues to characterize differential methylation, methylation entropy, and allele-specific methylation (ASM) associated with noncoding genetic variants. Genome-wide analysis identified extensive cancer-associated differentially methylated regions (DMRs), with hypermethylated DMRs significantly enriched near transcription start sites and transcriptional regulatory regions. Integration with transcriptomic datasets revealed strong inverse relationships between promoter methylation and gene expression, while 5-hydroxymethylcytosine (5hmC) levels positively correlated with transcriptional activity across gene bodies. Using fragment-level methylation patterns enabled by long-read sequencing, we further quantified methylation entropy incorporating both 5mCG and 5hmCG states. Cancer-hypermethylated DMRs exhibited markedly reduced entropy, consistent with clonal fixation of methylation states during tumor progression. Entropy profiling across chromatin annotations demonstrated maximal epigenetic heterogeneity at partially modified enhancer-associated regions. To investigate cis-regulatory genetic effects, we developed a simple ASM framework (nanoASM) that can partition sequencing reads by allelic state and identifies allele-specific DMRs directly from long-read data. Compared with conventional population-level mQTL analysis, ASM demonstrated substantially improved statistical efficiency by leveraging within-individual contrasts and reducing sample-level heterogeneity. Although germline single nucleotide polymorphisms (SNPs) were largely shared between normal and tumor tissues, ASM patterns differed substantially, with tumor-associated ASM regions displaying significantly larger genomic span and stronger allelic methylation differences. Comparative analysis with TCGA prostate mQTL and GTEx prostate eQTL datasets demonstrated substantial concordance between ASM directionality and downstream transcriptional effects, particularly for variants located within DMRs and near transcription start sites. At the IRX4 prostate cancer risk locus, ASM identified an androgen-responsive regulatory domain overlapping AR ChIP-seq and H3K27ac peaks, nominating rs6885084 as a candidate functional variant. At the PSCA locus, ASM anchored by rs4736369 was associated with allele-specific methylation, chromatin activation, transcript abundance, and isoform usage. Together, these findings establish nanopore-based ASM analysis as a powerful approach for resolving functional noncoding variants and their regulatory domains they control in prostate cancer.

9
Heterochromatin organization and liquid-liquid phase separation: it is not about if but about when

Romero, H.; Arroyo, M.; Zhadan, A.; Muzzopappa, F.; Zhang, H.; Qin, W.; Mahmoud, M.; Leonhardt, H.; Erdel, F.; Cardoso, M. C.

2026-06-07 cell biology 10.64898/2026.06.03.729812 medRxiv
Top 0.3%
1.7%
Show abstract

Heterochromatin is a membraneless compartment within the cell nucleus. In recent years, a controversy arose on whether heterochromatin organization is driven by liquid-liquid phase separation or not. While many heterochromatin proteins were shown to undergo liquid-liquid phase separation in vitro, other studies reported that this does not happen in cells. Here, we tested the ability of heterochromatin proteins to generate heterochromatin barrier compartments in cells. We found that, while several proteins (H1.0, H1.4, HP1alpha, HP1beta, Mbd1, Mbd2 and MeCP2) form barrier compartments in mouse and/or human cells this differs between cell types. In addition, not all compartments in the same cell form barriers. We established and experimentally validated a model that predicted the ability to form barrier compartments is dependent on the protein accumulation in heterochromatin followed by the competition between compartments for the nucleoplasm pool of the protein and resulted in larger size for the barrier compartments. These findings resolve the existing controversy and rationalize how in cells heterochromatin compartments form and compete to establish dynamic barriers to the entry and exit of its components. HighlightsHeterochromatin barrier formation differs between proteins, cell lines and heterochromatin compartments within the cell. Barrier formation depends on heterochromatin anchors, including ligands and other scaffolds. Barrier compartments are defined by their larger size and higher protein enrichment. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=83 SRC="FIGDIR/small/729812v1_ufig1.gif" ALT="Figure 1"> View larger version (28K): org.highwire.dtl.DTLVardef@a63021org.highwire.dtl.DTLVardef@a20362org.highwire.dtl.DTLVardef@8c2390org.highwire.dtl.DTLVardef@72dde3_HPS_FORMAT_FIGEXP M_FIG C_FIG

10
MethylBench: A comprehensive benchmark of DNA methylation profiling methods across diverse sequencing platforms

Laufer, L.; Gasparoni, G.; Hentrich, T.; Sofan, L.; Admard, J.; Buena-Atienza, E.; Pogoda, M.; Ossowski, S.; Casadei, N.; Riess, O.; Haack, T.; Buchert, R.; Schulze-Hentrich, J.

2026-04-30 genomics 10.64898/2026.04.28.721268 medRxiv
Top 0.4%
1.4%
Show abstract

BackgroundDNA methylation can be profiled using multiple technologies that vary in resolution, coverage and cost. Yet systematic benchmarks across these methods remain scarce. MethodsWe compared six widely used technologies -- Illumina EPIC array, TWIST, Whole-Genome Enzymatic Conversion (WGEC), Reduced Representation Bisulfite Sequencing (RRBS), long-read genome sequencing (LR-GS) with Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT) -- using Genome in a Bottle (GIAB) reference samples and ten samples derived of blood and fibroblast cultures of 5 individuals. We assessed CpG coverage, consistency of differentially methylated cytosine (DMC) detection and genomic annotation, with particular attention to overlapping signals across assays. ResultsDespite major differences in assay design, all technologies consistently identified DMCs enriched in promoter and intronic regions, highlighting these loci as robust hotspots of epigenetic variability. Annotation redundancy strongly influenced initial interpretations, with CpG island-related categories largely disappearing once annotations were collapsed to unique features. Sequencing-based methods (WGEC, TWIST, ONT) achieved the most comprehensive coverage, whereas EPIC arrays reproducibly captured promoter-associated differences despite limited scope. ONT sequencing enabled direct, long-read-based methylation profiling with phasing capability and showed strong concordance with short-read sequencing methods after coverage filtering, but required higher and more uniform coverage to achieve reproducible CpG-level agreement. PacBio methylation profiles showed a coverage-dependent discrepancy, with cross-platform concordance plateauing in GIAB samples despite high mean coverage, indicating residual technology-specific biases beyond simple coverage effects. ConclusionsCross-platform benchmarking yields coherent biological insights when coverage and annotation redundancies are carefully addressed. Practically, EPIC arrays remain valuable for promoter-focused cohort studies, WGEC and TWIST enable genome-wide discovery and ONT provides unique phasing and multimodal potential. This comparative framework can guide method selection and support more robust interpretation of DNA methylation data across diverse platforms.

11
A radial map of the budding yeast genome reveals novel organizational principles

Laenen, G.; Yip, W. H.; Baquero Perez, M.; Cournac, A.; Bienko, M.; Taddei, A.

2026-05-08 molecular biology 10.64898/2026.05.06.722996 medRxiv
Top 0.4%
1.1%
Show abstract

The eukaryotic genome is non-randomly organized within the nucleus, with positioning linked to function. Still, genome-wide radial maps are missing for the majority of experimental model systems. We adapted Genomic loci Positioning by Sequencing (GPSeq) to Saccharomyces cerevisiae, enabling high-resolution mapping along the nuclear center-periphery axis. GPSeq confirms known spatial features and shows that peripheral telomeres and centromeres impose long-range constraints extending up to 200 kb, restricting short chromosome arms from the nuclear interior. Telomere repositioning to the nuclear center, either artificially or during quiescence, reorganizes much of the genome through inward movement of sub telomeric regions and compensatory shifts of mid-arm chromatin outward. In quiescence, reduced centromere peripheral localization further alters genome organization. While transcription has a modest impact on radial positioning in all studied conditions, we uncover that in the absence of centromere or telomere constraints, GC-content functionally organizes chromatin in the nucleus. Graphical abstractThe budding yeast genome is spatially organized in a manner highly dependent on the positioning of centromeres (CENs) and telomeres (TELs). Anchoring of these chromosome landmarks constrains the positioning of adjacent chromatin up to 200 kb within the same radial zone. Beyond this range, genome organization is non-random, with processes like transcription and features such as GC- content associated with specific radial positions in the nucleus. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=142 SRC="FIGDIR/small/722996v1_ufig1.gif" ALT="Figure 1"> View larger version (29K): org.highwire.dtl.DTLVardef@843adborg.highwire.dtl.DTLVardef@1343eb3org.highwire.dtl.DTLVardef@1009b03org.highwire.dtl.DTLVardef@c10245_HPS_FORMAT_FIGEXP M_FIG C_FIG

12
CpG Atlas: A centralized multi-layer database and AI interface for DNA methylation research

Armstrong, J. F.; Wahi, S.; Borrus, D.; Sehgal, R.; Rizvi, S.; Zhang, S.; Jacques, M.; Eynon, N.; van Dijk, D.; Higgins-Chen, A.

2026-06-03 bioinformatics 10.64898/2026.05.30.729020 medRxiv
Top 0.4%
1.1%
Show abstract

DNA methylation research has vastly expanded over the past decade, producing a wealth of epigenome-wide association studies, biomarker algorithms such as epigenetic clocks, technical performance analyses, and functional annotations for CpG sites. However, these resources remain fragmented across dozens of databases and supplementary files within manuscripts, forcing researchers to spend time and effort on data cleaning and integration prior to meaningful analyses. No single resource currently unifies this information into a centralized, easy-to-query framework. Here, we present CpG Atlas, a curated relational database that integrates 18 distinct annotation layers encompassing over 1.2 million CpG sites across all four generations of Illumina methylation arrays (HM450K, EPIC v1, EPIC v2, and MSA). Built on a snowflake schema with a canonical probe identifier hub implemented in SQL, CpG Atlas consolidates over 800,000 CpG-trait associations, results from Mendelian randomization analyses, CpG membership across 81 epigenetic clocks, array manifest information, and probe reliability data. It further includes specialized layers such as solo-WCGW, CoRSIVs, PRC2 binding, transposon and retroelement annotations, tissue-specific differentially methylated positions across 17 tissues, and hallmarks of aging and cancer. To maximize utility and ease of use, the database is paired with an interactive web tool and a natural language-to-SQL query interface, enabling users to quickly perform complex multi-dimensional queries. Detailed documentation about every data source and table is also provided, facilitating the identification and interpretation of relevant studies. We demonstrate the utility of CpG Atlas through two case studies: a systematic enrichment analysis revealing distinct functional signatures across 16 epigenetic clocks, and an iterative biomarker discovery workflow for IBD that leverages cross-layer integration. Because it is readily scalable simply by adding or updating tables in the database, CpG Atlas provides a continuously evolving and extensible infrastructure for the epigenetics community that supports collaborative research, interpretable biomarker development, and integrative analyses across the growing landscape of epigenetic data.

13
Single molecule footprinting measures low nucleosome occupancy in mature spermatozoa of mice and men

Gaspa-Toneu, L.; Shi, H.; Ozonov, E. A.; Gill, M. E.; De Geyter, C.; Peters, A. H. F. M.

2026-07-01 genomics 10.64898/2026.06.30.735528 medRxiv
Top 0.4%
1.1%
Show abstract

Nucleosomes are fundamental units of DNA packaging and gene regulation in eukaryotes. In mammalian sperm, most nucleosomes are replaced by protamines causing extreme chromatin compaction. Various epigenomic studies reported conflicting results on the distribution of residual nucleosomes in mammalian sperm, questioning their potential role in mediating intergenerational inheritance of paternal epigenetic information. Here we performed single-molecule footprinting through Nucleosome Occupancy and Methylome (NOMe) sequencing and applied the Bayesian statistical model nomeR to determine frequencies of nucleosome removal and retention at 103 specific genomic regions in thousands of developing haploid spermatids and mature spermatozoa of mice. While we readily detected footprints of nucleosomes and the transcription factor CTCF in round spermatids, chromatin became transiently highly accessible in elongating spermatids with loss of such footprints, indicating extensive chromatin reprogramming during spermiogenesis. In mature sperm, following nuclear decondensation with recombinant nucleoplasmin, we measured nucleosome occupancy frequencies ranging ~1.2 to 1.7% at mouse loci. In human sperm, nucleosome occupancy varied between ~2.3 to 4.5% at 163 genomic loci profiled. Contrasting mice, chromatin in ~25% of human sperm was accessible upon reducing disulfide bonds between protamines arguing for species specific protamine packaging. Our findings support a stochastic rather than programmed potential role of residual nucleosomes in mammalian sperm in regulating paternal gene expression during ensuing embryonic development.

14
Explainable AI identifies H3K18ac as a new marker of active enhancers

Maqsood, K.; Polvora Brandao, D.; Wolfe, J.; Kayyar, B.; Clinciu, C. G.; Bosnea, R. A.; Grant, O. A.; Boulet, F.; Bell, C. G.; Ficz, G.; Madapura, P.; Hagras, H.; Zabet, N. R.

2026-06-11 genomics 10.64898/2026.06.09.731088 medRxiv
Top 0.4%
1.1%
Show abstract

Enhancers are non-coding regions of DNA that regulate gene transcription, yet the mechanisms underlying enhancer activity remain incompletely understood. Despite extensive experimental and computational efforts, we still lack accurate enhancer maps in many human cells, tissues and disease contexts. Here, we developed several Artificial Intelligence (AI) models (Convolutional Neural Networks (CNN), XGBoost, Logistic Regression (LR) and an eXplainable Artificial Intelligence type2 Fuzzy Logic based System (type2-FLS)) to predict enhancers across different human and mouse cell lines. While all models display high accuracy in the cell lines they were trained on, our results confirmed that type2-FLS, and, partially, CNN, LR and XGBoost perform consistently well in cell lines unseen during training, supporting the generalisation of the models. Most importantly, type2-FLS identified H3K18ac as an important enhancer mark along with many novel putative enhancers, which display the same epigenetic signatures as experimentally identified ones. We have validated some of these novel enhancers by both global epigenetic perturbations and directed enhancer epigenetic rewriting (CRISPRi). Interestingly, seven epigenetic marks in humans and five in mouse are sufficient to annotate enhancers without losing accuracy. Overall, we have deciphered the epigenetic code of mammalian enhancers and annotated enhancers in multiple human and mouse cell lines.

15
Quantifying the Information Capacity of DNA Methylation as an Epigenetic Memory System

De la Fuente, I. M.; Carrasco-Pujante, J.; Fedetz, M.; Legarreta, L.; Malaina, I.; Camino-Pontes, B.; Perez-Yarza, G.; Martinez, L.; Cortes, J. M.; Lopez, J. I.

2026-07-10 systems biology 10.64898/2026.06.28.735086 medRxiv
Top 0.4%
1.1%
Show abstract

The information content of the genome has been extensively analyzed. However, a comparable quantitative framework for DNA methylation is still lacking. Without such quantification, the magnitude of this regulatory and dynamic epigenetic structure remains conceptually imprecise, even though methylation dysregulation is strongly linked to disease-related phenotypes and altered cellular identity. Here we address this gap by applying Shannon information theory to DNA methylation. We first consider methylation marks as binary or probabilistic regulatory states and estimate the theoretical upper-bound information capacity of the human methylome under simplifying assumptions. We then progressively refine this estimate by incorporating biologically relevant constraints, including methylation bias, bimodal methylation distributions, local CpG correlation, genomic regulatory class, and cell-type-discriminative methylation patterns. This approach allows us to distinguish between theoretical methylation capacity, statistical methylation entropy, and biologically interpretable regulatory information. Finally, we consider methylation information from a discriminative perspective, analyzing its contribution to distinguishing cell types and regulatory cellular states. Within this framework, mutual information between methylation patterns and cell identity provides a biologically constrained estimate of methylations role as an epigenetic identity code. Our layered analysis reconciles megabit-scale methylome capacity with compact, biologically interpretable identity signatures. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=113 SRC="FIGDIR/small/735086v1_ufig1.gif" ALT="Figure 1"> View larger version (73K): org.highwire.dtl.DTLVardef@f0f0fdorg.highwire.dtl.DTLVardef@5d8a1eorg.highwire.dtl.DTLVardef@116debdorg.highwire.dtl.DTLVardef@79530e_HPS_FORMAT_FIGEXP M_FIG C_FIG

16
ZBTB38 requires an extended N-terminal zinc finger network to read mCpG- and discriminate TpG-containing DNA sequences

Boster, J.; Gangi, C.; Hudson, N. O.; Billings, D. E.; Guerra Castanaza Jenkins, B. L.; DIng, V. L.; Buck, B. A.

2026-05-07 biophysics 10.64898/2026.05.04.722763 medRxiv
Top 0.4%
1.0%
Show abstract

Methylation of cytosine bases in the CpG context (mCpG) is an essential regulatory mechanism cells use to spatially and temporally orchestrate access to genomic regions and mediate transcription. In many diseases, DNA methylation patterns become inappropriately distributed leading to aberrant transcriptional outcomes. Methyl-CpG binding proteins (MBPs) are key epigenetic mediators that selectively recognize mCpG sites, translating these signals into discrete transcriptional responses. ZBTB38 is a zinc finger (ZF) MBP that uniquely harbors two sets of five ZF clusters; each capable of selectively distinguishing mCpG sites. While the cognate DNA sequence and molecular basis for selective mCpG recognition have been defined for the ZBTB38 C-terminal (C-term) ZF domain, the molecular basis for differentiating DNA targets by the N-terminal (N-term) ZF domain remained uncharacterized. Here we report the mCpG-containing consensus sequence for the ZBTB38 N-term ZFs and demonstrate that unlike the other two ZBTB MBP family members ZBTB33 (Kaiso) and ZBTB4, the three shared core ZF domain discriminates against binding to TpG-containing DNA, and that at least one additional N-term ZF is required to stabilize DNA engagement. In addition, we demonstrate that each ZBTB38 ZF domain exhibits preferential target recognition for their respective cognate methylated DNA consensus motif. These findings expand understanding for how ZBTB38 differentially mediates epigenetic-based transcriptional process in normal and disease-state cells by providing new insight into the molecular basis by which the ZBTB38 N-term ZF domain differentiates DNA targets and offering further insight into the interplay between the N- and C-term ZF domains in directing cellular activities.

17
CREPAS: a reproducible nascent chromatin sequencing analysis pipeline for epigenome replication studies

Ruiz-Perez, S.; Du, Q.; Biran, A.; Groth, A.; Alcaraz, N.

2026-06-25 bioinformatics 10.64898/2026.06.21.732899 medRxiv
Top 0.4%
1.0%
Show abstract

Chromatin-based genomics data are essential for understanding genome regulation and the mechanisms underlying epigenetic memory. Recent methods such as ChOR-seq and SCAR-seq assess histone modifications and chromatin-associated proteins during and after replication, capturing chromatin states that contribute to memory across cell divisions. Current tools for chromatin data analysis lack scalability and reproducibility across computing infrastructures, offer limited parameters, and are applicable only to a few sequencing techniques, ignoring the information from nascent chromatin assays. To address these challenges, we developed CREPAS, a Nextflow pipeline for analyzing nascent and parental chromatin sequencing data, including ChIP-seq, ChOR-seq, SCAR-seq, OK-seq, ATAC-seq, CUT&RUN, and CUT&Tag, and derivative protocols. CREPAS provides an end-to-end solution, from quality control to advanced analyses, including downsampling, peak calling, annotation, and visualization. By harnessing quantitative assays such as qChIP-seq and qChOR-seq, the normalization methods in CREPAS allow to compare the restoration kinetics of individual marks or proteins across replication timepoints. Moreover, the pipeline includes calculations such as fork directionality and partitioning using OK-seq and SCAR-seq data, linking replication dynamics to epigenetic inheritance. CREPAS is a valuable resource that enhances the efficiency and reproducibility of nascent chromatin sequencing data analyses, enabling the study of chromatin replication and propagation of epigenetic states. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=80 SRC="FIGDIR/small/732899v1_ufig1.gif" ALT="Figure 1"> View larger version (24K): org.highwire.dtl.DTLVardef@29ad26org.highwire.dtl.DTLVardef@26b9acorg.highwire.dtl.DTLVardef@67dcb8org.highwire.dtl.DTLVardef@cbc299_HPS_FORMAT_FIGEXP M_FIG C_FIG

18
Quantifying the contribution of DNA conformational flexibility to transcription factor binding on nucleosomal DNA uncovers indirect readout across diverse TF families

Dey, U.; Martinez, G. S.; Kumar, R.; Yella, V. R.; Kumar, A.

2026-06-06 bioinformatics 10.1101/2025.05.21.655105 medRxiv
Top 0.5%
1.0%
Show abstract

BackgroundEukaryotic gene regulation depends on transcription factors (TFs) recognizing short DNA motifs within chromatin. Many of these motifs lie within nucleosomes, where DNA is sharply bent, rotationally phased, and constrained by histone-DNA contacts. Yet only a subset is occupied in any cellular context. Motif identity alone, therefore, cannot fully explain selective TF engagement with nucleosomal DNA. We asked whether sequence-derived DNA conformational flexibility provides an interpretable representation of sequence context relevant to TF recognition on nucleosomes. ResultsWe compiled five DNA flexibility descriptors in the Python package DNAflexpy, representing bendability, torsional deformation, backbone conformational variability, and stiffness. We built quantitative models of TF binding affinity across 226 datasets from a high-throughput in vitro TF-nucleosome binding assay. Flexibility-augmented models improved prediction over mononucleotide baselines in most datasets, with smaller but reproducible gains over trinucleotide baselines. The gains were not uniform: they varied across TF families and were concordant with DNA shape-fluctuation features, suggesting that DNAflexpy descriptors capture a sequence-encoded structural signal. In PIONEAR-seq data, model performance generalized across nucleosomal templates in a TF- and sequence-dependent manner. Beyond prediction, position-resolved flexibility footprints revealed deformation signatures at cognate motifs and flanking regions across diverse TF families. For SOX11, model-derived footprints aligned with DNA shape fluctuations from nanosecond-to-microsecond molecular dynamics trajectories of SOX11-bound nucleosomes, consistent with independently observed DNA conformational dynamics and bound-state stabilization. The in vivo data showed a similar but more context-dependent pattern. OCT4 occupancy tended to correlate with local flexibility, whereas GATA3-pioneered regions showed flexibility coupled with altered rotational positioning of cognate motifs. Flexibility-augmented classifiers further improved discrimination of occupied nucleosomal motifs across ENCODE datasets. Torsional flexibility features, particularly twist dispersion and trx, were most informative for classification. ConclusionsSequence-derived DNA conformational flexibility provides a quantitative and interpretable representation of sequence context in TF recognition on nucleosomes. By augmenting sequence with structural information, these models help quantify and interpret an indirect-readout contribution in which DNA deformation tendencies may complement motif sequence and DNA shape. This framework may help explain why only selected motif instances are engaged in chromatin, without treating flexibility as independent of primary sequence.

19
Systematic toxicological study of PFOS/PFOA co-exposure driving prostate cancer: Core target identification, TME immune remodeling, and combination drug prediction

PAN, J.; ZHANG, Y.; YANG, A.; JIANG, L.; SHEN, Y.; SUN, Y.; ZHU, J.; FAN, M.; SHI, J.

2026-05-12 pharmacology and toxicology 10.64898/2026.05.07.723528 medRxiv
Top 0.5%
1.0%
Show abstract

BackgroundPer- and polyfluoroalkyl substances (PFAS), particularly perfluorooctane sulfonate (PFOS) and perfluorooctanoic acid (PFOA), are persistent organic pollutants ubiquitous in the environment. Epidemiological evidence has closely linked them to an elevated risk of prostate cancer (PCa). However, the precise molecular mechanisms by which combined PFOS/PFOA exposure promotes prostate cancer and their dynamic effects on the tumor microenvironment remain unclear. MethodsThis study constructed a multi-module analytical framework integrating network pharmacology and computational biology: (1) Through ADMET toxicity prediction, multi-database target collection (three-way Venn analysis), panoramic GO/KEGG enrichment, focused androgen receptor (AR) axis analysis, GWAS genetic association validation, protein-protein interaction (PPI) network construction, machine learning-based independent screening, and a relaxed intersection strategy, we systematically identified PFOS/PFOA-prostate cancer core targets. (2) Subsequently, a PFAS-PTS score weighted purely by Cox coefficients was employed to drive gene set variation analysis (GSVA)-based pathway enrichment, tumor microenvironment (TME) deconvolution, ordinary differential equation (ODE)-based kinetic modeling, and drug intervention prediction. ResultsTarget collection identified 100 shared PFOS/PFOA-prostate cancer targets, from which 18 core targets were determined after multi-module screening. These targets were significantly enriched in the AR signaling axis, the PI3K-AKT pathway, and cell cycle regulation. Molecular docking confirmed strong binding affinities of PFOS/PFOA with AR (-9.49/-8.56 kcal/mol), AKT1 (-7.56/-6.93 kcal/mol), and PTEN (-6.36/-6.08 kcal/mol). GSVA revealed that the G2M checkpoint and E2F target gene pathways were significantly upregulated in the high-risk group (padj < 0.001), whereas the androgen response pathway was downregulated (padj = 4.8e-4). TME deconvolution (GSE141445, NNLS) revealed a significantly increased proportion of tumor cells (PCa) (p = 2.4e-4) and markedly reduced CD8+ T cell infiltration (p = 5.7e-4) in the high-risk group, indicating immunosuppressive microenvironment remodeling. ODE-based kinetic modeling confirmed that PFAS promoted tumor cell proliferation and suppressed immune surveillance in a dose-dependent manner. Drug intervention simulation demonstrated that the combination of enzalutamide and Alpelisib achieved optimal tumor cell inhibition (33.9% predicted by the ODE model). ConclusionPFOS/PFOA promote prostate cancer progression primarily through multi-target synergy involving AR axis disruption, PI3K-AKT pathway activation, and cell cycle dysregulation, while reshaping an immunosuppressive tumor microenvironment. The integrative computational framework established in this study provides systematic computational evidence for risk assessment and therapeutic intervention in PFAS-associated prostate cancer.

20
dCBP-mediated histone lactylation contributes to meiotic chromosome maintenance.

Nakayama, K.; Saito, D.; Hayashi, Y.

2026-05-18 developmental biology 10.64898/2026.05.15.725312 medRxiv
Top 0.5%
1.0%
Show abstract

Histone lactylation is a recently identified histone post-translational modification (PTM) that links energy metabolism to chromatin regulation. Although histone lactylation has been implicated in transcriptional activation, its function in meiotic chromatin remains unclear. Previously, we identified enrichment of multiple histone lactylation marks within the meiotic karyosome, a highly condensed and transcriptionally repressive chromatin structure formed in Drosophila oocytes. Here, through an RNAi-based screen, we identified the CBP family protein dCBP as a regulator of histone lactylation in the karyosome. Germline-specific knockdown of dCBP preferentially reduced histone lactylation, including H4K8 lactylation, and caused premature disruption of the synaptonemal complex, abnormal egg chamber development with excess nurse cells, reduced egg production, and decreased embryonic viability. Corresponding histone acetylation marks were comparatively less affected than histone lactylation by dCBP knockdown. Together, our findings provide evidence that dCBP-mediated histone lactylation contributes to meiotic chromosome maintenance and suggest a potential link between energy metabolism and meiotic chromatin regulation.